How to Design AI Evaluations You Can Actually Trust →
How to Write Reliable Rubrics for LLM-as-a-Judge Evaluations →
Join us for this episode of Firebase After Hours as we dive deep into the world of AI Evaluations!
Testing and measuring the success of AI agents and skills is one of the hardest parts of building reliable AI applications. In this livestream, we'll explore the concept of "Eval-Driven Development" and discuss how you can write, run, and interpret the results of your evals using open-source frameworks like Inspect AI.
We'll be sharing real-world case studies and behind-the-scenes insights from the Firebase team, including:
* Why evals matter and what it actually takes to run them.
* Specific examples of Eval-Driven Development in practice.
* Tips on designing evaluations and crafting effective scorers and rubrics.
* Common (anti-)patterns we've discovered while evaluating our own internal AI skills.
Whether you're just starting out with AI or looking to improve the reliability of your existing agent skills, tune in to learn how to measure what matters!
Subscribe to Firebase →
|
Google DeepMind's Paige Bailey catches u...
Andrew from the Flutter team catches up ...
Everyone is talking about AI agents and ...
Welcome to this comprehensive course on ...
How to Design AI Evaluations You Can Act...
[Sorry for the streaming hitches - we sh...
Starting in Flutter 3.47, the material_u...
Shoppers don’t follow a straight path an...
Festive season is when consumer demand t...
Save this for your next project with hea...